Data Quality (DQ) Processor using Unity Catalog

Calibo Accelerate supports end-to-end data quality management for data stored in Unity Catalog by integrating data ingestion, data profiler, data analyzer, data validator, and issue resolver into a structured pipeline.

To simplify data quality implementation and improve operational efficiency, data ingestion and data profiling are combined into a single Data Integration stage, while data analyzer, data validator and issue resolver are handled in a separate stage of data quality called DQ Processor.

In the data integration stage, data is ingested from source systems such as Microsoft SQL Server using Databricks for data integration and loaded into Databricks Unity Catalog. As part of the same stage, data profiling is performed to analyze structural and statistical characteristics of the ingested data, helping establish baseline quality metrics and metadata early in the pipeline.

The DQ Processor stage focuses on applying validation rules, identifying data quality issues, and managing issue resolver workflows based on the profiling results and defined constraints.

This topic describes how to create a data quality processor pipeline that reads data from Databricks Unity Catalog, applies data quality rules to validate and resolve identified issues, and writes the processed results back to Databricks Unity Catalog for downstream use.

To create a DQ Processor job, you must complete the following high-level steps:

  1. Create a DQ processor job giving it an appropriate name.

  2. Select the source table and data processing type.

  3. Enable Rule Configuration.

    1. Generate rules by running rule suggestions job and notify users via email.

    2. View and add rules.

    3. Test rules on actual data.

    4. View rule results.

    5. Add analyzer rules as required.

    6. Add reference and schema based on your use case.

    7. Click Run Validations.

    8. Check the validation results.

  4. Enable Issue Resolver

    1. Select Data Processing Mode:

      • Passed Validation Data

      • Full Data

    2. Select the constraints to run on the data.

  5. Select the target schema and tables to store successful records, Validator records and rejected Issue Resolver records.

The data quality processor pipeline has the following nodes:

Databricks Unity Catalog (data lake) > Databricks Unity Catalog (data quality node) - DQ Processor

Prerequisites

To create or run a data quality processor job using Unity Catalog, you must complete the following prerequisites:

  • Get access to a Unity Catalog data lake configuration listed under Configuration > Cloud Platform Tools & Technologies > Databases and Data Warehouses.

  • Ensure that the Statistics Common Table is configured in the Unity Catalog instance to store statistical data and metadata information.

  • Ensure that the Databricks cluster is configured with a Databricks Machine Learning Runtime before enabling anomaly detection.

To create a data quality processor job

  1. Sign in to the  Calibo Accelerate platform and navigate to Products.

  2. Select a product and feature. Click the Develop stage of the feature and navigate to Data Pipeline Studio.

  3. Add the Data Lake stage > Databricks Unity Catalog node, then configure the data lake node.

  4. Add the Data Quality stage > Databricks Unity Catalog > DQ Processor node.

  5. Connect the data lake and data quality nodes to each other.

  6. Click the data quality processor node and complete the following steps to create a data quality processor job:

 

Related Topics Link IconRecommended Topics What's next? Databricks Templatized Data Integration Jobs